fix(rocm): make CPU/Hybrid MoE graph replay safe - #378
Conversation
- Add hip_compat.h shim mapping CUDA runtime API to HIP equivalents - Update pinned_tensor.cpp to compile under both nvcc and hipcc - Add ROCm detection in arch.py (is_rocm, get_rocm_gfx_arch, is_gfx11xx_family) - Guard NVIDIA arch checks to return None on ROCm - Skip nvcc version check in _toolchain.py when on ROCm - Add ROCm build path in setup.py (ROCM_HOME, amdhip64, --offload-arch) - Add _hip_cflags() in kernel/utils.py for JIT compilation on ROCm - Add is_rocm() and driver_hip_version() in backend.py - Add rocm-smi fallback in __main__.py for clangd generation - Add TODO(ROCm) for NCCL->RCCL, flashinfer/sgl_kernel ROCm builds, Triton autotune RDNA3 tuning, PDL equivalent, hiprtc JIT cache - Add AMD ROCm classifier in pyproject.toml
Fail closed to eager execution when the HIP stream-memory handshake cannot survive capture and replay. Add a ROCm 7.14 graph batch-memop path with executor-owned signal and parameter storage, dynamically size graph flag slots, preserve the existing CUDA module API, and cover the safety and multi-format replay paths.
What: - Remove .agents/learnings and .plans/rocm-consolidation files from the branch. - Remove internal increment and plan-path references from source comments and public installation docs. - Keep implementation comments that explain correctness, ownership, profiler intent, source attribution, or ROCm safety behavior. - Clarify public ROCm documentation: gfx1100 has recorded serving smoke on ROCm 7.2.1; the ROCm 7.14.x container is a reference environment, and other target cells remain compile-only until physical serving evidence exists. Why: - Keep merge surface focused on code, tests, reproducibility tooling, and user-facing documentation. - Prevent private planning history, review workflow language, stale plan paths, and local process notes from entering the upstream repository. - Avoid presenting compile success or a reference container as cross-target serving or performance proof. Related upstream work informing this branch: - PR FlashML-org#132: portable ROCm/HIP foundation. - PR FlashML-org#133: TVM-FFI index/store portability. - PR FlashML-org#135: RCCL tensor-parallel communication. - PR FlashML-org#136: native GGUF ROCm build and Q4_0 kernels. - PR FlashML-org#137: earlier AMD serving bring-up. - PR FlashML-org#217: source-fork ROCm, Qwen3.5 GGUF, and performance experiments. - PR FlashML-org#241: gfx1150 build, JIT, Triton, and attention hardening. - PR FlashML-org#260: gfx1151 validation and fallback/build evidence. - PR FlashML-org#316: HIP graph-capture-safe expert copies. - PR FlashML-org#378: CPU/Hybrid MoE graph replay safety. - Local branch milestones: 436263f, 926c1e8, e1d1856, 8a70c7e, and e5fd30f. Evidence: - 170 focused tests passed after cleanup. - gfx1100 is the only target with end-to-end Qwen3.5 GGUF serving smoke recorded here. - Remaining matrix targets are compile-only; no new throughput claim is published without a matching A/B manifest.
|
Thanks for building this — we filed #350 and did not expect a fix, let alone one validated on the Two things below: a ROCm datapoint your HIP memop work bears on directly, and an offer. 1. Your HIP memop port has a second consumer, and it is the half that is still privateReading the diff, the thing that struck us is that the HIP #311's disk-streamed PLE table resolves its stream memops by
−27.3 % at 10k, −26.2 % at 100k, roughly 14 ms per token, against the −2.6 % … +0.2 % We can only bound the split, not measure it. Sampled from Where your PR lands on it: the eager half is now reachable — on ROCm your module-level And the ROCm-7.14 fact you documented in that file — that ordinary None of this is a request. It is a note that the two land in the same place. 2. An offer against your draft-exit criteriaYou list "reproduce the replay suite on RDNA3", "complete repeated graph rebuild/replay and We can help with parts of the middle one, on RDNA4 rather than RDNA3. We run two Radeon AI We can also exercise it at TP=2. To be exact about what is new there: we have reported TP=2 Three honest caveats before you count on any of it. We cannot give you the "long-running serving stability" half. We serve We cannot promise a date. The box is a shared machine with a live soak on it. Say what you It is a second gfx1201 datapoint, not a clean- No quality claim from us in any case — we have no fidelity instrument on this box, which is why Platform2 × AMD Radeon AI PRO R9700 ( Our tree is upstream Box labelling, since this box has two cards: every number in §1 — decode, TTFT, the 47.7 GiB |
Summary
Fixes #350, where CPU/Hybrid MoE produces silently incorrect output during ROCm graph replay while eager execution remains correct.
hipMallocSignalMemoryand explicit graph batch-memory-op nodesCpuMoeExecutor, so their lifetime covers every graph replay without module-global growthmemop_submit/memop_syncpath to limit NVIDIA regression riskDependency and scope
Depends on #132. This branch is based directly on the current #132 head (
c0713e4). #132 itself is unchanged.This is intentionally a Draft stacked PR while the native ROCm graph path receives RDNA3 and long-running model-serving stability validation. Until #132 merges, GitHub will also show the dependency commits in this PR; the diff will collapse to this single follow-up commit after the base lands.
CUDA and ROCm retain the same externally visible CPU-MoE ordering and results. Only their GPU/CPU synchronization implementations differ:
cuStreamWriteValue64/cuStreamWaitValue64wrappersValidation
git diff --checkpassesDraft exit criteria